1554 stories
·
0 followers

Thousands will die if EPA repeals power plant climate rules, lawsuit says

1 Share

Days after finalizing a rule that would eliminate climate rules for power plants, the Environmental Protection Agency (EPA) was sued by groups who say the EPA’s repeal was shockingly short-sighted and risked leaving the country’s single largest source of industrial climate pollution unchecked.

On Thursday, the American Lung Association, the American Public Health Association (APHA), the Clean Air Council, Clean Wisconsin, the Environmental Defense Fund, and the Natural Resources Defense Council (NRDC) filed a petition requesting that the US Court of Appeals for the DC Circuit review if the EPA’s action conflicts with the Clean Air Act.

“Clean air is a basic human right,” APHA CEO Georges C. Benjamin said.

A physician and long-time public health policy leader, Benjamin warned that the EPA was “weakening” “life-saving standards” and stripping away “vital protections for all, and especially for children, pregnant people and communities already overburdened by pollution.” The lawsuit alleges that the rule change could cause thousands of premature deaths.

In the EPA’s announcement of the proposed repeal, the agency claimed that its action was based on “the best reading of the Clean Air Act.” Benjamin disagrees.

“Power plant pollution threatens the health of millions of Americans and fuels climate change, worsening extreme heat, poor air quality, and other serious health risks,” Benjamin said. He urged the court to recognize that the “EPA must maintain strong clean air and climate standards so every community can live a healthy life.”

The EPA may struggle to defend the rule given that its core mission is "to protect human health and the environment." Meredith Hankins, NRDC’s federal climate legal director, accused the EPA of abdicating its legal responsibilities. She appeared confident that the EPA repeal will not survive court scrutiny.

“The Clean Air Act and Supreme Court precedent demand that the EPA address climate pollution from the largest industrial source in the nation,” Hankins said. “The EPA’s legal reasoning is fatally flawed, so we are going to court.”

"Devastating consequences" of EPA's rollback

So far, the EPA has not acknowledged any potential harms of the rule change. Instead, the agency proposed rescinding all remaining greenhouse gas emissions requirements for power plants, claiming that the restrictions produce “virtually no benefits.”

Rather than meaningfully addressing the risks that advocacy groups describe, the EPA relied on an argument that anyone who’s ever monitored a climate lawsuit will recognize: that climate change is global and “public health harms are too uncertain, conjectural, remote, and convoluted to tie specifically to the US power sector.”

Overall, the agency appears more focused on claiming to have delivered $310 billion in savings on utilities for Americans, while saying that the repeal is “expected to unleash the full potential of America’s vast energy resources, including coal and natural gas.” The announcement also said the power plants will “save an additional $370 million in direct compliance costs.”

Advocacy groups say the EPA is turning a blind eye to a wide range of harms that could be easily addressed by maintaining the status quo. Lawrence Hafetz, the legal director of the Pennsylvania-based Clean Air Council, warned that the “EPA is senselessly discarding readily attainable carbon standards for power plants mere weeks after the UN showed the world on a trajectory to exceed the heating limit needed to avoid severe levels of economic, ecological, and social disruption.”

Impacts will be felt throughout the US, groups forecasted. For example, Brett Korte, a senior staff attorney for Clean Wisconsin, said he is bracing for “devastating consequences” if the rollback is allowed to stand.

“Here in Wisconsin, we’re experiencing warmer winters with less snow and ice cover, more exceedingly hot days in summer, and more frequent extreme weather events, like flooding and tornadoes, due to climate change,” Korte said. "On top of that, Wisconsin has the ninth-highest death rate from particulate matter pollution from fossil fuel combustion. EPA abandoning its responsibility to regulate carbon pollution would only make these problems worse.”

The EPA’s timing is terrible, and communities will experience unnecessary deaths if the court does not intervene, said Vickie Patton, who serves as general counsel for the Environmental Defense Fund.

“More Americans than ever are suffering from the dangerous heat and more powerful storms caused by that pollution and are facing the ensuing skyrocketing insurance bills and health costs,” Patton said. By increasing power plant climate pollution in this moment, the EPA rule change “will cause thousands of premature deaths and billions of dollars in additional health costs,” she predicted.

“We are taking action in the court of law to ensure [the] EPA carries out its responsibilities under our nation’s clean air laws to protect the American people from this harmful pollution and to address the health risks and financial costs for people across the country,” Patton said.

Thousands of Americans have submitted comments on the rule change so far, and the EPA continues to seek public input. At the end of this month, the EPA will hold a public hearing on the proposed repeal, an a public comment period extends through October.

Read full article

Comments



Read the whole story
Share this story
Delete

EPA seeks to eliminate remaining greenhouse-gas rules for power plants

1 Share

On the heels of the hottest summer on record and a United Nations reportwarning that global temperatures are likely to push climate risks to “increasingly dangerous heights,” the Trump administration overturned most of a Biden-era rule limiting climate pollution from power plants, the second-largest source of greenhouse gas emissions.

The Environmental Protection Agency announced a final plan Monday that guts the 2024 Carbon Pollution Standards, which the agency said exceeds its authority under the Clean Air Act by requiring control technologies “that are not adequately demonstrated.” The agency also proposed revoking “all remaining GHG emissions requirements for power plants,” arguing their emissions “have no material impact on climate change.”

The plan to revoke the Carbon Pollution Standards was initially revealed last spring.

In 2007, the Supreme Court ruled that the Clean Air Act’s definition of air pollutants includes greenhouse gases, which cause climate change, and that if the EPA concludes that GHGs pose a danger to human health and the environment, it must act to reduce those emissions. The EPA made that “endangerment finding” in 2009.

“The Trump administration is now kicking the legs out from under the entire legal framework for regulating climate pollution by arguing that climate pollution does not harm human health or welfare,” said Zealan Hoover, a former senior advisor to the EPA during the Biden administration.

Hoover and other environmental and climate experts have not yet had a chance to review the specific arguments the agency is making to justify its actions.

Maggie Coulter, senior attorney at the Center for Biological Diversity’s Climate Law Institute, expects the Trump administration to make a “non-endangerment finding” so they don’t have to regulate greenhouse gas emissions at all. “That goes against really well-established science,” Coulter said. “Will this stand up in court? We would argue, definitely not.”

Experts on the fossil fuel industry’s decades-long attempts to deny the reality or consequences of climate change see the same tactics in the Trump administration’s latest actions. President Donald Trump called climate changethe “greatest con job ever perpetrated on the world” last year during a speech at the United Nations.

“Who needs Big Oil to deny climate science, undermine science-based decision-making and endanger peoples’ lives and livelihoods when the U.S. government will do it for them?” asked Geoffrey Supran, director of the Climate Accountability Lab at the University of Miami.

“The government can choose to deny reality,” said Ben Franta, an associate professor of climate litigation at the University of Oxford, who has highlightedthe role of false solutions in delaying action on climate change. “But the real damages from global warming will continue to accumulate, from heat waves and fires to superstorms and floods, and calls for accountability will only increase.”

Since returning to office, Trump has overseen a multi-pronged attack on the legal and regulatory framework to protect public health and the environment from the climate-warming consequences of greenhouse gas emissions.

In February, the EPA overturned the endangerment finding that underpinned its authority to regulate greenhouse gases from motor vehicles, the largest sourceof climate pollution. The same month, the agency attacked states’ rights to pass stricter climate regulations by using a parliamentary trick to revoke decades-old Clean Air Act waivers that underpinned California’s vehicle emissions standards, though a federal judge has temporarily blocked that move. And now, in going after the Carbon Pollution Standards, the agency aims to remove limits on power plants.

Part one was repealing the endangerment finding, Hoover said. Now that that’s done, they’re working on part two, rolling back different regulations that relied on the endangerment finding, and part three, which would prevent future efforts to regulate climate pollution.

The first Trump administration repealed a bunch of public health regulations, then Trump lost his re-election bid, and the Biden administration reinstated many of them, he said. Now, they are not just repealing them “but burning them to the ground,” in an effort to make it incredibly hard for a future administration to come back in and re-regulate public health pollution, Hoover said. “It’s really quite pernicious.”

EPA spokespersons did not respond to a request for comment about what science they are using to support the claim that greenhouse gas emissions don’t contribute significantly to dangerous air pollution.

The climate and health benefits of the Carbon Pollution Standards “substantially outweigh the compliance costs,” Biden’s EPA announced in June 2024. The agency pointed to a regulatory impact analysis that projected reductions of 1.38 billion metric tons of carbon pollution through 2047, the equivalent of preventing nearly a year of emissions from the entire U.S. electric power sector or avoiding the annual emissions of 328 million gas-powered cars.

The entire regulatory framework is based on cost-benefit analyses, Hoover said. But so far, the Trump administration is “heavily discounting, if not completely refusing” to quantify public health benefits in regulations, he said.

In announcing the power plant regulations repeal, Trump’s EPA said only that by unleashing coal and natural gas, its new rule will save $310 billion and its proposal to rescind “every remaining greenhouse gas standard for the power sector” would save an additional $370 million in direct compliance costs. It also said the rules would save American families and businesses billions more.

Independent studies, however, have reported that coal power is driving higher utility costs for consumers. And elevated emissions of toxic air contaminants triggered by the repeal would lead to increases in health damages costing up to $476 billion, according to a cost-benefit analysis by the nonprofit group Resources for the Future.c

The two largest sources of climate emissions in the United States are the power and transportation sectors, Hoover said. “And with today’s regulatory rollback, the Trump administration has now finished its work to end the climate regulations on both.”

Liza Gross is a reporter for Inside Climate News based in Northern California. She is the author of The Science Writers’ Investigative Reporting Handbook and a contributor to The Science Writers’ Handbook, both funded by National Association of Science Writers’ Peggy Girshman Idea Grants. She has long covered science, conservation, agriculture, public and environmental health and justice with a focus on the misuse of science for private gain. Prior to joining ICN, she worked as a part-time magazine editor for the open-access journal PLOS Biology, a reporter for the Food & Environment Reporting Network and produced freelance stories for numerous national outlets, including The New York Times, The Washington Post, Discover and Mother Jones. Her work has won awards from the Association of Health Care Journalists, American Society of Journalists and Authors, Society of Professional Journalists NorCal and Association of Food Journalists.

This story originally appeared on Inside Climate News.

Read full article

Comments



Read the whole story
Share this story
Delete

Hackers reveal how Flock cameras really track cars and people

1 Share

Hackers ripped down a Flock camera above a roadway, made a near-complete copy of the data stored inside it, and shared the files with 404 Media and WIRED, revealing in new detail how exactly Flock Safety’s cameras track the movements of both vehicles and people. The hackers say they are also publishing details on how they managed to obtain the software, in the hopes that other people may copy them.

The breach provides an unprecedented look inside a system that Flock has described as protected by on-device encryption. The hackers were able to copy the camera’s storage and recover an encryption key stored on the device, which unlocked videos of thousands of vehicle detections. The hackers shared the material with 404 Media and the transparency nonprofit Distributed Denial of Secrets, which shared the data with WIRED. 404 Media and WIRED then analyzed those files as part of a joint investigation.

While much of the automatic license plate reader’s most sensitive storage remained encrypted and inaccessible, the joint analysis of the recovered data shows that software running on the device explicitly detects people as well as vehicles, license plates, and bicycles. The camera can produce dozens of images of a single passing vehicle and, according to several weeks of recovered logs, generated more than a million images. Its computer-vision software also sometimes isolated bumper stickers and other graphics, including, in one case, an American flag patch on a motorcyclist’s saddlebag.

The act of removing the camera and dumping its software shows that some people are not content with just destroying or removing the cameras. Across the country, multiple people have been arrested for allegedly tampering with or otherwise sabotaging Flock’s cameras. In response, some towns have announced that they are going to stop using Flock’s cameras altogether, and in one case, a police department even made a fake, 3D-printed Flock camera case in order to bait potential vandals.

“Why just destroy them when we can reverse engineer them and find the secrets of those spying on us?” one of the hackers, from a collective calling itself stegan0gram, said in an interview. “We liberated hardware in the field, disarmed them, and proceeded with reverse engineering of the cameras and associated solar equipment.”

Flock’s cameras photograph passing vehicles and send the images and other data to the company’s servers. There, Flock’s system presumably reads the license plate and can identify characteristics, such as the vehicle’s color, make, and model. Flock then makes these time-stamped records searchable by whichever local agency owns or has access to the cameras. But in many cases, Flock’s system also allows other police departments from all over the country to search those cameras as part of the company’s national network. In Alpharetta, Georgia, for example, WIRED found that records from the city’s Flock cameras were accessible to more than 2,000 agencies, including police departments, colleges, airports, and, inexplicably, the Office of Inspector General for the federal General Services Administration.

This national network has been a selling point for Flock but also a deep source of controversy. 404 Media revealed that local cops were performing lookups in the national network on behalf of Immigration and Customs Enforcement, including in areas that banned working with immigration authorities or transferring license plate data out of state. 404 Media also revealed that a cop in Texas searched Flock cameras nationwide for a woman who self-administered an abortion. Those stories, among others, triggered a national conversation about whether people want Flock cameras, or automatic license plate readers more generally, in their communities.

And in the case of stegan0gram, the answer is clearly no.

The hackers said they were able to access the Android system on the camera and found two partitions—sections of its hard drive, essentially. A few of these were unencrypted, the hackers said, including one called “vendor” and another called “media.” The latter contained an encryption key that unlocked another part, which contained much of the media—the videos and stills—the camera took.

In early 2025, security researcher Jon “GainSec” Gaines reverse-engineered a Flock license-plate reader and documented flaws that could be used to gain root-level access. After Gaines disclosed his findings, the company acknowledged the findings but downplayed their severity, writing that the flaws required physical access to the device and that even someone who gained access to a camera “would still not be able to gain access to footage,” because images remained on the device only briefly after being transmitted to the cloud.

404 Media and WIRED analyzed the camera’s contents. The device’s processor is similar to those used in midrange smartphones, and it runs about 20 Flock-built apps that handle everything from detecting motion and taking pictures to classifying objects, uploading data, and receiving remote updates.

According to the code, when something moves into view, the camera takes a rapid series of photos. A typical passing vehicle generated about 28 images, though some produced more than 100. The camera uses different exposures to capture both the license plate and the wider scene, then scans the images, selects and crops useful frames, and sends them with other data to Flock over the cellular network. The camera itself does not appear to read the plate or identify the vehicle’s make, model, and color. That appears to happen on Flock’s servers.

According to our analysis, the camera’s logs recorded about 21 days of activity across several periods. During those windows, the device photographed roughly 50,200 vehicles and generated about 1.6 million images. On a typical day, it logged around 3,300 vehicles, with a high of 4,454. Those figures would vary considerably depending on where a camera is installed and how much traffic passes in front of it. The camera was almost certainly operating outside those periods, but older logs had been overwritten or were no longer recoverable from the device.

The software running on the camera explicitly detects people, something that is typically overlooked in discussions around Flock cameras. When it spots a person, it records where they appear in the image and how confident it is in the detection.

To test what the software could actually see, WIRED extracted the models from the camera’s files and ran them against test images and footage recovered from the device. The models readily detected people, including a selfie of a reporter. WIRED then ran them across 27,321 short videoclips stored on the camera. The clips were MP4 files, each about one to two seconds long, recorded at 1,024 by 768 pixels without audio. They were separate from the rapid bursts of higher-resolution still images the camera also takes as vehicles pass. The models detected people in 11 of the clips, all of them riding motorcycles. The small number is likely due to the camera’s position above a roadway, pointed down at passing traffic where pedestrians were unlikely to appear.

The tests also showed how broadly the camera’s license plate detector could interpret what it saw. In some cases it mistook bumper stickers, dealership frames, and other graphics for license plates and cropped them out as if they were plates. In one video of a passing motorcycle, the detector cropped an American flag patch on the rider’s saddlebag as if it were a plate.

Flock insists its cameras do not perform face recognition. WIRED and 404 Media found no evidence of any face-recognition capabilities in the camera’s software beyond ones included by default in the Android operating system. Those capabilities did not appear to be enabled or in active use.

In August, WIRED obtained frontend code for Flock’s police software, now called OS Investigate and previously known as Nightshift, and reconstructed portions of the tool. That software showed how Flock can use the records generated by its cameras, along with police files and commercial data, to identify drivers and surface vehicles that repeatedly travel together and search for people based on patterns of movement. The data provides a view of the other end of a system.

A Flock spokesperson said in a statement: “The unauthorized removal and tampering of a Flock camera is illegal.” When asked specifically about the encryption key stored on the camera, the company added, “Flock takes security seriously and maintains a public Vulnerability Disclosure Policy for security researchers to report potential vulnerabilities directly to us. We received no report through that process, and based on the limited information provided, we do not have enough detail to assess the claims being made. If the individuals identified legitimate vulnerabilities, we encourage them to submit their technical findings through our vulnerability reporting process so our security team can review them and take any appropriate action.”

One of the hackers said, “Being investigated is a legit concern and something we are trying to avoid. I'm sure our actions have attracted some attention as it is, but we are careful and try to keep a low profile.”

Noel Pichardo, a former Pawtucket, Rhode Island, police officer who became an outspoken critic of Flock after challenging his department’s use of the cameras, says he understands the activists’ frustration but worries that sabotaging devices could ultimately strengthen the case for them. “I think that type of vigilantism will only crystallize the police and the state at large in their belief that this tool is necessary,” Pichardo says. “The longer the state continues to ignore the groanings of their constituents who are against this type of surveillance, the more this will happen.”

The camera’s logs also show the camera struggling with storage. Its logs recorded more than 27,000 “no space left on device” errors while trying to save full-resolution images, along with tens of thousands of related errors, crashes, and reboots. At the same time, about every two minutes, code checked that the camera was still running and logged the message, “Who’s a good boy?!” More than 12,000 of those messages appear in the recovered logs.

When the camera did restart, another service left a final message in the logs: “A reboot was requested! ¡Adiós, Amigos!”

This story originally appeared on wired.com.

Read full article

Comments



Read the whole story
Share this story
Delete

Microsoft exec called AI scraping the “largest theft of labor in human history”

1 Share

For years, Microsoft and OpenAI have fought to keep certain information out of the public eye in their fight with news organizations that have accused the AI firms of teaming up to violate copyright laws by stealing tons of news content to train AI.

However, now the details that should never have been marked confidential are starting to leak. In a motion for summary judgment that was unsealed Thursday from news plaintiffs led by The New York Times, internal documents are exposed that news groups alleged show exactly how Microsoft and OpenAI viewed the threat to news before unleashing new AI products like ChatGPT and Copilot.

Perhaps most explosively, Microsoft Director of Applied Science Brent Hecht repeatedly warned in documents that scraping news for AI training was “an astonishing theft of unprecedented proportions,” calling it perhaps the “largest theft of labor in human history,” news orgs said. In another document, Hecht contradicted Microsoft and OpenAI’s argument that training AI on news content is fair use, suggesting that the plan to widely scrape news made “a complete mockery of the idea of ‘fair use.’”

Over at OpenAI, ChatGPT head Nick Turley wrote in an internal message that publishers would face an “existential threat” from commercial products trained on news content that can be used to substitute news providers. One Microsoft document even described a “doom loop,” news orgs said, “that will hurt the performance of our models and the entire web at the same time.”

“It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain,’” that document said.

Data from both firms shows that this prediction was accurate. Microsoft recorded 83–93 percent drops in click-through rates for some news plaintiffs, and 51–94 percent drops for others. Add to that reporting on low click-through rates from ChatGPT search results and news organizations’ own reporting on traffic declines. Suddenly, it becomes easier to see how declining news revenue could ultimately rob chatbots of the abundant streams of reliable information that supposedly makes them such groundbreaking tools.

Meanwhile, “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use,” Hecht acknowledged in a Microsoft document.

News organizations say they’re ready to go to trial because there’s so much “compelling evidence of substitution.” If they can prove that chatbots are replacing them in their own markets, while serving to spit out excerpts of articles verbatim, they think that one-two punch may eviscerate Microsoft and OpenAI’s fair use arguments.

“The future not just of journalism but of responsible AI too depends on preserving incentives for humans to produce the creative works on which a healthy society depends,” news groups argued.

Chatbots are “largely substitutive, period”

Under oath, Microsoft CEO Satya Nadella testified that AI companies shouldn’t be violating news sites’ terms of use by dodging paywalls. But over at OpenAI, internal messages showed that when a staffer, Nick Ryder, informed President Greg Brockman that “a hack” was found for OpenAI crawlers “to get around” the NYT paywall, Brockman replied, “Ah, nice.”

Nadella also acknowledged that chatbots have served as substitutes for news platforms, describing the chatbot as stealing clicks from news sites by “giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”

There’s consensus on that at OpenAI, where a software engineer said in an internal message that “no matter how prominently we show the links, users won’t click.”

OpenAI’s Turley agreed that there is “no good reason to click” when the chatbot provides information, the motion said. He also seemingly suggested that the doom loop was already in motion, describing chatbots as “largely substitutive, period” and predicting that they “will get more and more substitutive as they get better.”

News groups argued that insiders' own statements should be damning.

“With respect to outputs that are substantially similar to training or grounding sources, courts have rejected claims that copying news articles to provide a product that substitutes for demand for news is fair use,” news groups argued.

Microsoft disclaims exec's comments

News groups tried many different tactics to test if Microsoft and OpenAI products would output their news articles verbatim. Their motion shows they went further than early strategies where they would ask chatbots to provide access to entire news stories by repeatedly asking “what’s the next line?”

In some cases, news organizations found that chatbots would generate long excerpts of articles when users requested summaries of articles. Other flagged outputs were generated by asking for key bullet points of articles. Particularly successful were prompts requesting that chatbots “rate the bias” of news articles. Chatbots also reproduced portions of articles if users asked them to pick any article off a certain site’s homepage.

In their motion, news plaintiffs have only asked the court to rule on infringed articles where outputs “demonstrate extensive verbatim overlap,” because they’re confident that the “substitutive purposes of defendants’ copying weigh against fair use.” Legal concerns with other articles will be raised at trial, they said.

OpenAI did not immediately respond to Ars’ request to comment.

However, a Microsoft spokesperson defended Microsoft’s AI products as a transformative fair use that don’t substitute for news sites. The spokesperson said that Nadella’s testimony touched on “broad principles and changes underway in how people find and consume information,” which were merely “observations” that “should not be confused with conclusions about copyright questions before the Court, which Microsoft addresses in its filings.”

Regarding Hecht’s comments, the spokesperson claimed that those documents only “reflect one employee’s individual perspective, are not a legal analysis, and do not represent the company’s views.”

Steven Lieberman, counsel for the New York Daily News and seven of its sister papers, disagrees. He told Ars that “the evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong.”

“Throughout this case Defendants insisted that these documents be treated as confidential so that the public could not see them,” Lieberman said. “Well, now the cat is out of the bag. Finally, the world can see what OpenAI and Microsoft thought all along about the fairness of their own behavior.”

Microsoft exec described "accidental cover up"

News plaintiffs have argued that regardless of the individual expressing the views, the internal documents make clear that firms anticipated that verbatim outputs would harm news sites. Further, they alleged that instead of preventing the outputs, the firms tried to make it harder for news groups to test chatbots by creating a filter that Hecht suggested could be perceived as an “accidental cover up” because it would result in “people who have a right over the content having less visibility into what was used for training."

News groups are also upset that instead of listening to insiders warning that scraping news was theft, Microsoft and OpenAI never chose to license content, allegedly usurping them in another market in ways they couldn't anticipate.

Specifically, their motion accused Microsoft of violating “industry norms” by selling a dataset purchased for Bing as training data for OpenAI, allegedly doing so without consulting news groups that would not have approved of that repurposing of their consent to basic search engine crawling. Further, OpenAI allegedly “acted improperly” by obtaining a NYT dataset with 1.8 million articles from a third party that was bound to an agreement that the data wouldn’t be used for commercial purposes. OpenAI’s employees knew it “would not be appropriate” to use that data “to train a model,” but they did it anyway, news groups alleged.

For news groups, the problem isn’t just Microsoft and OpenAI, but all the AI firms that are following their lead in "free-riding" on their content, the motion said. Most notably, after ChatGPT’s launch, Google’s AI Overviews was quickly introduced and started absorbing even more traffic that previously went to news sites.

If courts don’t clarify that AI firms must license news content, both news publishers and AI firms could be doomed, news plaintiffs argued. One Microsoft internal document agreed that “there is a ‘real risk’ that GenAI could ‘significantly disrupt’” the “employment of the very people who generated the data on which the foundation model was trained,” they noted. Microsoft even included a cartoon illustrating the problem of LLMs destroying their own supply chains, they said:

Cartoon in a Microsoft internal document. Credit: via News Plaintiffs

“AI companies remain powerless to break out of this ‘doom loop,’ because, while the industry as a whole would benefit if every company paid to sustain the continued production of the creative works their technology depends on, each individual company is better off taking content for free while others pay,” news groups argued.

As evidence of this blind greed, their motion emphasized that Brockman wrote that he was “deeply motivated by the gazillions” that could be gained by commercializing OpenAI’s technology.

“Finding that copying news for AI is not fair use would solve this prisoners’ dilemma by putting all AI companies, OpenAI and Microsoft included, on an even footing,” news organizations said.

This story was updated with a quote from New York Daily News counsel Steven Lieberman. 

Read full article

Comments



Read the whole story
Share this story
Delete

Kondex: A declaration indexing tool

1 Share

It’s been a very long time since I’ve written on this blog, but wanted to pick this up again, and thought tools would be an excellent topic.

Like most people in the Software Industry, I’ve been spending increasing amounts of time getting comfortable with LLM-based tools. This includes not only using tools such as GitHub Copilot and Claude Code, but also figuring out better ways to make our engineering team used to them and able to use them effectively.

In the next few days, I’d like to cover some of the internal tools I’ve built to help support this journey.

Note: Right now, these tools are purely internal and not published externally. However, if someone finds them useful, let me know and I’ll see what I can do.

First, I’d like to cover a tool called kondex.

What is Kondex?

Kondex is a utility tool that aims to simplify how LLM tools navigate a code base. There’s plenty of tools out there that claim to do this by indexing the code itself and then providing search operations on top of the index.

Kondex aims to solve the same issue, but from a slightly different perspective. Kondex does not apply any kind of RAG or vector-driven search. Instead, it works by building an index of all declarations (like classes, methods and other members), and then providing primitives to search through these declarations.

Why build it?

This is an excellent question. Like I said, there’s a number of similar tools out there (many open source), but after testing several of them, I couldn’t find one that would do exactly what I wanted and worked well with our specific requirements.

The biggest issues I ran into with other tools were:

  • Size of the code base: Our connector team works off a fairly big Git monorepo from which we build and deliver 1000s of artifacts, and contains millions of lines of code. Many of the tools I tried would need well over 10 minutes to even do the initial indexing, and even incremental indexing wasn’t very performant.
  • Support for Git worktrees: Our team relies heavily on using Git Worktrees. Most of the tools I tried are completely unaware of worktrees, and treat each separate branch on disk as a completely different repo, requiring generating a full index every time you created a new worktree. This was unacceptable for our use case.
  • Ignoring files: Our team makes heavy use of code generation. Our build process not only generates a lot of code files, but also copies stuff around while doing builds, all of which goes into a ./Release folder that’s excluded by .gitignore. The problem is most tools don’t know anything about this, and will happily index all source files on disk, polluting the index with either duplicate files or just generated code that’s irrelevant.

Solving these issues was the primary driver behind developing kondex.

Architecture

Kondex is implemented using Go. I wanted to use a language that would produce native tools for best performance. I am not an expert at Go, but that’s where Claude Code becomes very useful. It helped me build the initial working prototype fast and evolve it over time to fix issues and support new features. The following diagram shows the basic architecture:

kondex architecture

Parsing source files is achieved through the use of Tree Sitter, which does an excellent job of abstracting the complexity of building a syntax tree; kondex uses this as the base for its declaration extraction engine. Currently, kondex supports parsing declarations from the following languages:

  • Java
  • Kotlin
  • C#
  • Go
  • Python
  • TypeScript
  • Terraform

Because of Tree Sitter, adding support for other languages is substantially easier now. Note: kondex also supports parsing our own metadata definition files, which is done through a custom linear scanner instead of tree-sitter.

The actual index is stored in SQLite.

Building the Index

I mentioned before one of the key problems I wanted to solve with kondex was the cost of building a full index of a repository. Solving this required doing a few tricks that ended up being a very interesting part of the project. The following figure illustrates the process of doing an initial scan for a code base:

kondex initial scan

The most obvious initial solution for speed was parallelizing the work: kondex will indeed scan input files in parallel through the use of multiple threads. However, it has a single thread writing to the SQLite database and it can be the bottlleneck. This is the reason we chose to rebuild indexes at the end of the scan, instead of ensuring they are kept up-to-date as the data is written.

Supporting Git Worktrees required getting a bit more creative. Instead of deciding the list of files to scan using the file system, kondex takes a bit of a different approach:

  1. git ls-files -s gives us the list of all tracked files in the current worktree alongside its blob SHA.
  2. Indexed declarations are stored per blob SHA, and shared by every worktree or branch.
  3. Each worktree has it’s own snapshot, which tracks the path -> blob list

This means that when switching branches or worktrees, kondex is able to:

  • Keep a single index across branches or worktrees
  • Quickly figure out the files that have been changed, removed, or added from the current index and do an incremental scan, which can be done by comparing the new blob SHAs against the list already stored in the index.

Let’s illustrate this with an example. First, let’s build a full index of our current major branch:

> kondex scan
repo   C:/dev/src/<path>  (worktree: C:/dev/src/v26)
index  <path>\index.db
enumerated 145139 tracked files, 55834 indexable, 33173 under excluded directories, 46664 unique blobs (465ms)
cache: 0 hits, 46664 to read (116ms)
suspended query indexes and checkpoints for the load; both are done once at the end
  6162/46664 blobs, 77.0 MB, 129607 declarations
  12533/46664 blobs, 163.3 MB, 265517 declarations
  19370/46664 blobs, 260.2 MB, 410976 declarations
  26076/46664 blobs, 347.3 MB, 546584 declarations
  29517/46664 blobs, 399.7 MB, 615042 declarations
  37074/46664 blobs, 490.4 MB, 768798 declarations
  44792/46664 blobs, 590.0 MB, 923175 declarations
rebuilt query indexes (1.294s)

scanned v26 at c0049c880
  tracked    145,139 files (55,834 indexed)
  blobs      46,664 unique; 0 already cached, 46,597 read
  cache hit  0.0%
  read       618.1 MB
  extracted  963,379 declarations from 46,597 files
  syntax     326 files had error regions (0.70%); entities may be incomplete
  skipped    67 over size limit
  excluded   33,173 files under infoscripts/
  elapsed    18.9s  (enumerate 465ms, diff 116ms, read 15.29s, index 1.29s, write 1.18s, checkpoint 559ms)
  writer     busy 11.85s of the 15.29s read (78%), 2.58s of it committing

You can see doing a full index of over 45K files was done in less than 20 seconds. This is good enough that we can use kondex even without forcing the developer to manually trigger a full scan the first time the tool is used!

Now let’s switch to a worktree tied to a different major branch, and do an incremental scan:

> kondex scan
repo   C:/dev/src/<path>  (worktree: C:/dev/src/v25)
index  <path>\index.db
enumerated 147991 tracked files, 54909 indexable, 35470 under excluded directories, 41853 unique blobs (497ms)
cache: 34322 hits, 7531 to read (91ms)
suspended query indexes and checkpoints for the load; both are done once at the end
  2886/7531 blobs, 77.5 MB, 70259 declarations
  5915/7531 blobs, 177.1 MB, 150643 declarations
rebuilt query indexes (1.71s)

scanned v25 at 521d30840
  tracked    147,991 files (54,909 indexed)
  blobs      41,853 unique; 34,322 already cached, 7,493 read
  cache hit  82.0%
  read       225.5 MB
  extracted  192,292 declarations from 7,493 files
  syntax     55 files had error regions (0.73%); entities may be incomplete
  skipped    38 over size limit
  excluded   35,470 files under infoscripts/
  elapsed    9.85s  (enumerate 497ms, diff 91ms, read 5.63s, index 1.71s, write 1.12s, checkpoint 810ms)
  writer     busy 4.14s of the 5.63s read (73%), 714ms of it committing

Now we see the incremental scan for a different (older) branch being done in less than 10 seconds. This is a fairly high value due to the number of modified files, but branches with fewer changes will see very fast scans.

What’s next?

I will continue going over kondex functionality in a later post, by showing what you can do once an index has been built.

Read the whole story
Share this story
Delete

Kondex: Searching the index

1 Share

In my last post, I introduced kondex, a declaration indexing tool I wrote for our team. Today, I’d like to spend some time going over how to use the index to navigate the code base.

Note: kondex will automatically detect a stale the index on any command and do an incremental scan if necessary.

The first entry point to the index is the search command, which uses a case-insensitive substring match on the declaration name and finds matching declarations:

> kondex search OnViewportWidthChanged
method  Winterdom.Viasfora.Rainbow.RainbowLines#OnViewportWidthChanged(object,EventArgs)  src/Viasfora.Rainbow/RainbowLines.cs:154-157
method  Winterdom.Viasfora.Text.CurrentLineAdornment#OnViewportWidthChanged(object,EventArgs)  src/Viasfora.Core/Text/CurrentLineAdornment.cs:83-85
method  Winterdom.Viasfora.Text.PresentationMode#OnViewportWidthChanged(object,EventArgs)  src/Viasfora.Core/Text/PresentationMode.cs:28-31
method  Winterdom.Viasfora.Util.ToolTipWindow#OnViewportWidthChanged(object,EventArgs)  src/Viasfora.Rainbow/Util/ToolTipWindow.cs:76-81

4 results

Note: The result of a search in kondex lists what we call a ref, which will usually take the form of Type#member(parameter types), with constructors using the <init> name.

Members declared outside of a type will be owned by the module/package they are declared in:

  • Go functions will be owned by their package (e.g. scan#Run())
  • Python / TypeScript functions at module level will have an empty owner (e.g. #UseThis())

When searching, exact names will rank first, then prefixes, then any other matches. A term containing ., # or ( is matched against the full ref rather than the name, so search 'Rainbow#On' works.

There are a couple of interesting additional filters that can be used for search. First, you can narrow down results to only declarations of a specific kind:

> kondex search --kind class Rainbow
30 results for "Rainbow" (kind class); more matched

class  Winterdom.Viasfora.Rainbow.Rainbows                          src/Viasfora.Rainbow/Rainbows.cs:4-19
class  Winterdom.Viasfora.Tags.RainbowTag                           src/Viasfora.Core/Tags/RainbowTag.cs:6-11
class  Winterdom.Viasfora.Rainbow.RainbowLines                      src/Viasfora.Rainbow/RainbowLines.cs:52-359
...
class  Winterdom.Viasfora.Rainbow.RainbowKeyProcessorProvider       src/Viasfora.Rainbow/RainbowKeyProcessor.cs:11-24

30 results shown; more matched — raise --limit, or narrow with --kind/--path

Note that in the previous example, kondex found over 30 matches, so the result is partial. You can do a search --count to get just the total number of results without any extra details.

You can also limit results by only searching within a specified path:

> kondex search Rainbow --path .\src\Viasfora.Core\
class        Winterdom.Viasfora.Tags.RainbowTag                              src/Viasfora.Core/Tags/RainbowTag.cs:6-11
constructor  Winterdom.Viasfora.Tags.RainbowTag#<init>(IClassificationType)  src/Viasfora.Core/Tags/RainbowTag.cs:8-10
field        Winterdom.Viasfora.Guids#RainbowOptions                         src/Viasfora.Core/Guids.cs:9
field        Winterdom.Viasfora.PkgCmdIdList#cmdidRainbowNext                src/Viasfora.Core/PkgCmdIdList.cs:15
field        Winterdom.Viasfora.PkgCmdIdList#cmdidRainbowPrevious            src/Viasfora.Core/PkgCmdIdList.cs:14

5 results

Outline

Another extremely useful feature of kondex is being able to return an outline for a file (meaning, a simplified view of the declarations in the file). Again, this saves LLM tools from having to read the file and extract them directly:

> kondex outline .\src\Viasfora.Languages\Sql.cs
src/Viasfora.Languages/Sql.cs  (11 declarations)

  class        Sql                             8-25
  field          #knownContentTypes            10-11
  property       #SupportedContentTypes        12
  property       #Settings                     13
  constructor    #<init>(ITypedSettingsStore)  15-21
  method         #NewBraceScanner()            23-24
  class        SqlSettings                     27-43
  property       #ControlFlowDefaults          28-32
  property       #LinqDefaults                 33-35
  property       #VisibilityDefaults           36-38
  constructor    #<init>(ITypedSettingsStore)  40-42

The outline command also accepts a folder path as an input, in which case it will produce an outline of all files in the directory (recursively). Since this could be a very large result, it will only report a number of findings (200 declarations by default). You can use --limit and --offset to narrow it down.

Members

One common question an LLM might want to answer is: What members does this type declare? That’s what the members command is for:

> kondex members Winterdom.Viasfora.Rainbow.CharPos
struct Winterdom.Viasfora.Rainbow.CharPos  src/Viasfora.Languages/CharPos.cs
  public struct CharPos

  field        CharPos#ch                    5
               private readonly char ch
  field        CharPos#state                 6
               private readonly int state
  field        CharPos#position              7
               private readonly int position
  field        CharPos#Empty                 8
               public static CharPos Empty
  property     CharPos#Char                  10
               public char Char
  property     CharPos#State                 11
               public int State
  property     CharPos#Position              12
               public int Position
  constructor  CharPos#<init>(char,int)      14-15
               public CharPos(char ch, int pos) : this(ch, pos, 0)
  constructor  CharPos#<init>(char,int,int)  17-21
               public CharPos(char ch, int pos, int state)
  method       CharPos#ToString()            27-29
               public override string ToString()

10 members

There are two options that can be very useful to members:

  • The --public flag will ask kondex to only list public members of the specific class. The rules for each language will determine what is considered public:
    • Java / C# includes only strict public members (not protected, internal, or package)
    • Go will define it based on capitalization
    • TypeScript will define it based on export
  • The --inherited flag will ask kondex to not only list members directly defined in the specified type, but also in the superclasses (or interfaces) in the source (it will obviously not return anything that comes from external, unindexed libraries). Since kondex only does a basic lexical scan, the matching is done by name, which means this is supported strictly as a best-effort.

Show

Once you’ve found the declaration you are looking for, you can show it:

> kondex show "Winterdom.Viasfora.Util.ToolTipWindow#OnViewportWidthChanged(object,EventArgs)"
method Winterdom.Viasfora.Util.ToolTipWindow#OnViewportWidthChanged(object,EventArgs)
  file         src/Viasfora.Rainbow/Util/ToolTipWindow.cs:76-81
  declaration  private void OnViewportWidthChanged(object sender, EventArgs e)
  declared in  Winterdom.Viasfora.Util.ToolTipWindow

Note how this gives you the complete method declaration, along with a location (file + line range). This is the ideal case when all you care is about what the declaration looks like (for example, you just want the method signature)

One interesting feature of show is that kondex is aware of language-specific syntaxes for documenting declarations, such as javadoc comments or C#’s /// comments, and will return those as part of the show command where available:

> kondex show "Winterdom.Viasfora.Options.MainOptionsControl#Dispose"
method Winterdom.Viasfora.Options.MainOptionsControl#Dispose(bool)
  file         src/Viasfora/Options/MainOptionsControl.Designer.cs:12-17
  declaration  protected override void Dispose(bool disposing)
  declared in  Winterdom.Viasfora.Options.MainOptionsControl

  Clean up any resources being used.

  @param disposing true if managed resources should be disposed; otherwise, false.

Note: The above example is interesting in that it searches the declaration based on a partial prefix: The ref used is missing the parameter list!

There are cases when an LLM might want more than a signature. For example, it might need to read what the method implementation looks like. In that case, kondex show also has the --body argument, which gives you its source or body:

> kondex show "Winterdom.Viasfora.Util.ToolTipWindow#OnViewportWidthChanged(object,EventArgs)" --body
method Winterdom.Viasfora.Util.ToolTipWindow#OnViewportWidthChanged(object,EventArgs)
  file         src/Viasfora.Rainbow/Util/ToolTipWindow.cs:76-81
  declaration  private void OnViewportWidthChanged(object sender, EventArgs e)
  declared in  Winterdom.Viasfora.Util.ToolTipWindow

  76      private void OnViewportWidthChanged(object sender, EventArgs e) {
  77        this.tipView.ViewportWidthChanged -= this.OnViewportWidthChanged;
  78        if ( this.tipView.ViewportRight > this.tipView.ViewportLeft ) {
  79          this.ScrollIntoView(this.pointToDisplay);
  80        }
  81      }

This is very useful to LLM tools as it saves them the need to open the file directly and scan the line range.

Tip: In some cases, you might have multiple declarations with the same name. In that case, you will want to either specify --path or --kind to ensure you get the right one. Note that if the input to show is ambiguous, it will provide a list to disambiguate:

> kondex show OnViewportWidthChanged
"OnViewportWidthChanged" matches 4 declarations:
  method  Winterdom.Viasfora.Rainbow.RainbowLines#OnViewportWidthChanged(object,EventArgs)  src/Viasfora.Rainbow/RainbowLines.cs:154-157
  method  Winterdom.Viasfora.Text.CurrentLineAdornment#OnViewportWidthChanged(object,EventArgs)  src/Viasfora.Core/Text/CurrentLineAdornment.cs:83-85
  method  Winterdom.Viasfora.Text.PresentationMode#OnViewportWidthChanged(object,EventArgs)  src/Viasfora.Core/Text/PresentationMode.cs:28-31
  method  Winterdom.Viasfora.Util.ToolTipWindow#OnViewportWidthChanged(object,EventArgs)  src/Viasfora.Rainbow/Util/ToolTipWindow.cs:76-81
retry with one of the refs above, or narrow with --kind/--path

Terraform

I thought including a short example of using kondex on code bases including terraform scripts would be interesting, as the syntax used can be somewhat different from other languages. For example, searching:

> kondex search dns
30 results for "dns"; more matched

data_source  data.azurerm_resource_group.dns-rg                  terraform/dev/dns.tf:29-32
data_source  data.azurerm_resource_group.dns-rg                  terraform/modules/cdatakube_assets/external-dns.tf:5-8
data_source  data.azurerm_resource_group.dns-rg                  terraform/prod/security.tf:65-67
...

30 results shown; more matched — raise --limit, or narrow with --kind/--path

You can see Terraform introduces its own set of declaration kinds (data_source, variable, resource, module, output, local, provider). Showing a declaration still works the same way:

> kondex show azuread_application.azuredns-sp
resource azuread_application.azuredns-sp
  file         terraform/prod/security.tf:28-31
  declaration  resource "azuread_application" "azuredns-sp"
  members      none

  Create Azure AD App for azure dns

An interesting aspect for Terraform is that refs use the exact same syntax the configuration language uses to refer to things.

The workflow

By now, you can probably guess there’s an implicit workflow to using kondex:

  • search to find the name
  • outline or members to see the shape
  • show when you want the complete declaration, show --body when you need the implementation

In other words, the whole purpose is to provide a tool that helps an LLM narrow down as much as possible the possibilities before it reads anything.

One thing worth being honest about: The examples above are almost all from the Viasfora code base, which, let’s face it, is very small. The numbers don’t reflect much complexity. However, the tool can have substantial impact when operating on a large code base, particularly one that has lots of larger source files. This is where kondex really shines, as it lets the LLM avoid having to read larger files directly, or spend multiple cycles trying to figure out the exact boundaries of a declaration for a narrow read.

In the next post, I will cover something that was hinted at in the original article, but not explicitly mentioned: Why kondex comes with a Claude Code plugin!

Read the whole story
Share this story
Delete
Next Page of Stories